The re-identification risk of Canadians from longitudinal demographics

نویسندگان

  • Khaled El Emam
  • David L. Buckeridge
  • Robyn Tamblyn
  • Angelica Neisa
  • Elizabeth Jonker
  • Aman Verma
چکیده

BACKGROUND The public is less willing to allow their personal health information to be disclosed for research purposes if they do not trust researchers and how researchers manage their data. However, the public is more comfortable with their data being used for research if the risk of re-identification is low. There are few studies on the risk of re-identification of Canadians from their basic demographics, and no studies on their risk from their longitudinal data. Our objective was to estimate the risk of re-identification from the basic cross-sectional and longitudinal demographics of Canadians. METHODS Uniqueness is a common measure of re-identification risk. Demographic data on a 25% random sample of the population of Montreal were analyzed to estimate population uniqueness on postal code, date of birth, and gender as well as their generalizations, for periods ranging from 1 year to 11 years. RESULTS Almost 98% of the population was unique on full postal code, date of birth and gender: these three variables are effectively a unique identifier for Montrealers. Uniqueness increased for longitudinal data. Considerable generalization was required to reach acceptably low uniqueness levels, especially for longitudinal data. Detailed guidelines and disclosure policies on how to ensure that the re-identification risk is low are provided. CONCLUSIONS A large percentage of Montreal residents are unique on basic demographics. For non-longitudinal data sets, the three character postal code, gender, and month/year of birth represent sufficiently low re-identification risk. Data custodians need to generalize their demographic information further for longitudinal data sets.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Reliability of Identification of Seniors at Risk Screening Tool in Predicting Functional and Mental Decline in Discharged Elderly Patients

Introduction: Emergency wards today are facing with an increasing numbers of older patients. Therefore, it seems important and essential to develop a short screening tool with an acceptable predictive power to identify the seniors being discharged from hospital and mean while are at risk of a decline in physical and mental performance, and thus, facing re-admission emergency wards in hospitals....

متن کامل

Never too old for anonymity: a statistical standard for demographic data sharing via the HIPAA Privacy Rule

OBJECTIVE Healthcare organizations must de-identify patient records before sharing data. Many organizations rely on the Safe Harbor Standard of the HIPAA Privacy Rule, which enumerates 18 identifiers that must be suppressed (eg, ages over 89). An alternative model in the Privacy Rule, known as the Statistical Standard, can facilitate the sharing of more detailed data, but is rarely applied beca...

متن کامل

Assessment of Facial and Cranial Development in Shirvanian Kurmanj Population Based on the Mean Biometric Factors from Birth to Maturity Age

Purpose: The aim of this study was to determine cranial & facial anthropometric Ratios and assessment of cranial & facial development in Shirvanian kurmanj population. Materials and Methods: This cross sectional analytical study was conducted randomly on 137 boys from shirvan, with normal face patterns. Facial and cranial ratios was estimated and compared. Data were analyzed by SPSS software. T...

متن کامل

Intelligent identification of vehicle’s dynamics based on local model network

This paper proposes an intelligent approach for dynamic identification of the vehicles. The proposed approach is based on the data-driven identification and uses a high-performance local model network (LMN) for estimation of the vehicle’s longitudinal velocity, lateral acceleration and yaw rate. The proposed LMN requires no pre-defined standard vehicle model and uses measurement data to identif...

متن کامل

De-identification Methods for Open Health Data: The Case of the Heritage Health Prize Claims Dataset

BACKGROUND There are many benefits to open datasets. However, privacy concerns have hampered the widespread creation of open health data. There is a dearth of documented methods and case studies for the creation of public-use health data. We describe a new methodology for creating a longitudinal public health dataset in the context of the Heritage Health Prize (HHP). The HHP is a global data mi...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:

دوره 11  شماره 

صفحات  -

تاریخ انتشار 2011